YWT Data Home

Category

Data Literacy

14 articles

Documentation Half-Life: How Federal Datasets Outlive the Records That Make Them Usable

Documentation Half-Life: How Federal Datasets Outlive the Records That Make Them Usable

Federal datasets often remain publicly accessible for decades after the internal documentation that explains them has quietly disappeared. When data dictionaries erode, variable definitions drift, and methodology notes vanish from agency servers, researchers are left reconstructing intent from incomplete evidence. This article examines how documentation decay happens, where it has already caused measurable harm, and what a practical metadata audit looks like before you commit to a dataset.

Precision Theater: How the Systematic Omission of Margin of Error Is Misleading America's Decision-Makers

Precision Theater: How the Systematic Omission of Margin of Error Is Misleading America's Decision-Makers

When data consumers encounter a statistic without its accompanying margin of error, they are not receiving incomplete information — they are receiving a distorted picture engineered, however unintentionally, to project false certainty. This article examines how the routine suppression of uncertainty disclosures in high-profile data reporting constitutes a form of statistical misinformation, and what data professionals can do to push back.

Hidden in Plain Sight: How Denominator Choices Are Distorting America's Most Cited Statistics

Hidden in Plain Sight: How Denominator Choices Are Distorting America's Most Cited Statistics

Per capita figures appear in nearly every policy brief, news headline, and research summary produced in the United States — yet the population base used to construct them is almost never disclosed or defended. The choice of denominator is not a technical footnote; it is a substantive analytical decision that can invert conclusions, obscure disparities, and quietly steer billions of dollars in resource allocation. Data professionals who fail to interrogate that choice are inheriting someone else'

Shrinking Samples, Expanding Claims: The Quiet Crisis of Attrition in America's Longitudinal Surveys

Shrinking Samples, Expanding Claims: The Quiet Crisis of Attrition in America's Longitudinal Surveys

Participation rates in the United States' most influential longitudinal surveys have been declining for decades, and the people who remain in these studies are increasingly unlike those who dropped out. The result is a compounding distortion that widens with every successive wave — yet the research claims built on this eroding foundation rarely acknowledge how much the ground has shifted beneath them.

The Invisible Overhead: Unpacking the True Cost of Reproducing Published Research

The Invisible Overhead: Unpacking the True Cost of Reproducing Published Research

Reproducing a published study is rarely as straightforward as its methods section implies. Across economics, public health, and social science, data professionals are absorbing enormous hidden costs in time, labor, and institutional goodwill—costs that rarely appear in any budget line but accumulate into a systemic crisis for the research enterprise. This article examines where those costs originate, who bears them, and what a genuinely replication-ready research standard would require.

Joined at the Seam, Broken at the Core: The Hidden Costs of Linking Federal Datasets

Joined at the Seam, Broken at the Core: The Hidden Costs of Linking Federal Datasets

Merging datasets across federal agencies can appear straightforward in a script but catastrophic in its consequences. Incompatible geographies, misaligned time windows, and conflicting unit definitions routinely corrupt merged results before a single statistical test is run. This article examines the structural fault lines that make cross-agency data linkage one of the most underappreciated sources of research error in the United States.

Temporal Validity and the Quiet Obsolescence of Data: A Framework for Assessing When Your Dataset Has Outlived Its Usefulness

Temporal Validity and the Quiet Obsolescence of Data: A Framework for Assessing When Your Dataset Has Outlived Its Usefulness

Researchers routinely treat historical datasets as perpetually reliable evidence, yet the real-world conditions those datasets once captured may have shifted beyond recognition. Economic disruptions, demographic realignments, and policy overhauls can render even widely cited datasets scientifically misleading long before anyone thinks to question them. This article examines the mechanisms of data decay and introduces a structured approach for evaluating freshness risk before a dataset is committ

The Provenance Problem: How Synthetic Training Data Is Quietly Corrupting the Scientific Record

The Provenance Problem: How Synthetic Training Data Is Quietly Corrupting the Scientific Record

Synthetic datasets — generated by AI models rather than collected from observed phenomena — are entering research pipelines at an accelerating rate, often without adequate documentation of their origins or limitations. When these records are treated as equivalent to empirically gathered data, the resulting findings carry a form of hidden uncertainty that standard peer review processes are not designed to detect. The implications for reproducibility, policy, and scientific credibility deserve urg

Borrowed Assumptions: What Pre-Packaged Datasets Are Doing to Your Research Before You Run a Single Query

Borrowed Assumptions: What Pre-Packaged Datasets Are Doing to Your Research Before You Run a Single Query

When researchers reach for a curated, ready-to-use dataset, they are not simply saving time — they are inheriting every undocumented decision made by whoever cleaned it first. This article examines how preprocessing choices baked into widely used datasets in economics, public health, and housing can silently distort findings. We offer a structured provenance auditing framework any data professional can apply before trusting a packaged source.

Beyond the Headline Numbers: Unlocking the ACS Tables Most Data Professionals Have Never Opened

Beyond the Headline Numbers: Unlocking the ACS Tables Most Data Professionals Have Never Opened

The American Community Survey publishes thousands of detailed cross-tabulated tables that go far beyond the population and income figures most researchers cite. From commute duration by occupation to housing cost burden by citizenship status, these overlooked tables offer analytical depth that headline ACS figures simply cannot provide. This guide walks through eight high-value table series and shows exactly how to retrieve them.

What Isn't There: How Systematically Absent Health Data Distorts US Research and Policy

What Isn't There: How Systematically Absent Health Data Distorts US Research and Policy

Missing data in major US health datasets is rarely random—and that distinction carries enormous consequences for researchers, policymakers, and the populations those policies serve. This article examines the structural patterns behind absent records in CDC and CMS data, illustrates how those absences have skewed published findings, and provides practical detection frameworks for data professionals working with incomplete health records.

Garbage In, Policy Out: Auditing the Structural Flaws in America's Most Trusted Federal Datasets

Garbage In, Policy Out: Auditing the Structural Flaws in America's Most Trusted Federal Datasets

Federal datasets like the American Community Survey and the Behavioral Risk Factor Surveillance System are foundational to US research pipelines — but embedded collection inconsistencies and demographic blind spots can silently corrupt downstream analysis. Before your next download, here is what critical data professionals need to examine. Understanding these structural limitations is no longer optional; it is a prerequisite for responsible research.